Tag
18 articles
Google launches Gemini 3.8 Live, a cost-effective alternative to OpenAI's GPT-Live-1, offering competitive audio AI performance at a fraction of the price.
OpenAI releases GPT-Live-1, bringing natural, full-duplex voice conversations to the API with enhanced instruction following and telephony support.
OpenAI releases GPT-Live-1 API, enabling developers to build apps with real-time, full-duplex voice interaction. The API scores 80.1% in interactivity tests, though it comes at a cost of $0.05 per minute.
Smallest.ai raises $13M to build ultra-fast voice AI that sounds genuinely human, aiming to create AI phone calls that can pass the Turing test.
PolyAI introduces Dialog-RSN-1, an audio-native dialog model that processes caller audio directly, fusing turn-taking, speech recognition, function calling, and response generation into a single system.
Fish Audio, a Palo Alto voice AI startup, has raised $52 million in a seed round by offering its best models for free and charging for latency and integration services.
A flurry of major funding rounds in AI, quantum computing, and deep-tech highlights growing investor confidence in transformative technologies.
ElevenLabs is reportedly in early talks for a $22 billion tender offer, nearly double its February funding round valuation, as the voice AI startup seeks to reward employees and attract more capital.
Coval raises $28M to stress-test AI voice agents before they reach real callers, applying safety methodologies from autonomous vehicles to ensure reliable performance.
A new open-source voice model named Audio Interaction listens continuously and makes real-time decisions every 0.4 seconds about when to speak or stay silent.
OpenAI has launched new voice intelligence features in its API, enhancing voice recognition accuracy and enabling more natural conversations. The updates have applications across customer service, education, and creator platforms.
Mistral AI's new TTS model, Voxtral, tackles the 'expressivity gap' in voice AI by combining autoregressive and flow-matching techniques for more emotionally expressive, multilingual speech synthesis.